Papers by Anmol Reddy Mekala
Alternate Preference Optimization for Unlearning Factual Knowledge in Large Language Models (2025.coling-main)
Copied to clipboard
Anmol Reddy Mekala, Vineeth Dorna, Shreya Dubey, Abhishek Lalwani, David Koleczek, Mukund Rungta, Sadid A. Hasan, Elita A.A Lobo
| Challenge: | Existing methods for large language models rely on negative feedback to suppress responses related to the forget set, which often results in nonsensical or inconsistent outputs, diminishing model utility and posing potential privacy risks. |
| Approach: | They propose an approach which combines negative feedback with in-domain positive feedback on the forget set and introduces new evaluation metrics to assess the quality of responses related to the forget sets. |
| Outcome: | The proposed approach avoids undesirable model behaviors while maintaining overall model performance. |